Papers with large-scale evaluation dataset

2 papers
TurnBack: A Geospatial Route Cognition Benchmark for Large Language Models through Reverse Route (2025.emnlp-main)

Copied to clipboard

Challenge: Existing studies on large language models have limited evaluation of their geospatial cognition . a unified framework for evaluating geospcial cognition in LLMs remains absent .
Approach: They propose a benchmark to evaluate the geospatial route cognition of Large Language Models . they propose 'pathbuilder' tool for converting natural language instructions into navigation routes .
Outcome: The proposed framework and metrics evaluate 9 state-of-the-art LLMs on route reversal task.
PERSONA: A Reproducible Testbed for Pluralistic Alignment (2025.coling-main)

Copied to clipboard

Challenge: Currently, preference optimization approaches fail to capture the plurality of user opinions . Currently used methods do not account for the pluralities of users and difference of opinion .
Approach: They propose a reproducible test bed to evaluate pluralistic alignment of language models . they generate user profiles from census data and use a large-scale evaluation dataset .
Outcome: The proposed model improves pluralistic alignment of language models with diverse user values . it generates a large-scale evaluation dataset with 317,200 feedback pairs .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations